Back

National Science Review

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match National Science Review's content profile, based on 21 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Engineered probiotic Escherichia coli-mediated intestinal nicotine clearance alleviates nonalcoholic steatohepatitis in mice

Zuo, N.; Cai, X.; Wang, W.; Ren, Z.; Jiang, Z.; Jiang, W.; Song, X.; Gu, Y.

2026-07-09 synthetic biology 10.64898/2026.07.02.736048 medRxiv
Top 0.2%
1.2%
Show abstract

Nicotine accumulates in the gut and drives non-alcoholic steatohepatitis (NASH) via the gut-liver axis, yet no effective clinical intervention is currently available. To address this challenge, the probiotic Escherichia coli Nissle 1917 (EcN) was engineered for in situ nicotine clearance in the gut. Mutational screening of nicotine oxidoreductase 2 (PpNicA2) identified a highly active variant, PpNicA2A107R. Its incorporation into EcN together with an electron transfer protein (CycN) and a newly identified transporter (T3/T7) yielded 80% nicotine-degrading activity. Chromosomal integration of this module generated a stable strain, EcN-N12, which in NASH mouse models depleted intestinal nicotine, rescued hepatic lipid metabolism, alleviated tissue damage, and intercepted the nicotine-mediated gut-liver axis pathological progression. This work thus offers an effective and clinically translatable approach for nicotine-associated diseases.

2
Complete elucidation and heterologous reconstruction of the biosynthetic pathway of camptothecin

Zhang, T.; Xiong, Y.; Chen, K.; Wu, S.; Yan, X.; Zhou, J.; Wang, Y.; Yang, C.; Wang, P.; Zhou, Z.

2026-07-08 synthetic biology 10.64898/2026.06.23.733941 medRxiv
Top 0.2%
1.2%
Show abstract

Camptothecin derivatives are first-line anticancer drugs used worldwide for the treatment of diverse malignant tumors. However, the biosynthetic pathway of camptothecin has remained elusive for five decades. Here, we fully map its entire biosynthetic route. We discovered five key missing enzymes (OpCAR, OpSDR11, OpCS, OpGH1, and OpSTR) via the combination of MALDI mass spectrometry imaging, single-cell RNA sequencing and co-expression analysis. Meanwhile, we demonstrated a free flavin mononucleotide triggered the non-enzymatic 6-5-6 to 6-6-5 fused-ring skeleton rearrangement, filling the last gap in camptothecin biosynthesis. Finally, we validated this identified pathway and achieved the de novo biosynthesis of camptothecin in Saccharomyces cerevisiae. These discoveries uncover the long-standing mystery underlying camptothecin and pave the way for manufacturing camptothecin and its derivatives through synthetic biology approaches.

3
NR4A3 knockdown ameliorates metabolic dysfunction-associated steatotic liver disease through ATF3 transcriptional repression

Liao, H.; Qin, B.; Zhou, L.

2026-06-30 pathology 10.64898/2026.06.24.734361 medRxiv
Top 0.2%
1.0%
Show abstract

Objectives; The role of nuclear receptor subfamily 4, group A, member 3 (NR4A3) in hepatic steatosis, inflammation, and insulin resistance (IR) within the context of metabolic dysfunction-associated steatotic liver disease (MASLD) remains largely underexplored. Consequently, this study aimed to examine NR4A3's impact on MASLD and the potential underlying mechanisms. Methods; We aimed to elucidate the functional role of NR4A3 in MASLD through its knockdown in cell culture and animal models. To establish the cell culture model of MASLD, LO2 cells were treated with free fatty acids (FFAs), while male C57BL/6 mice were fed a high-fat diet (HFD) to create the animal model. NR4A3 knockdown was achieved using specific short hairpin RNA (NR4A3-shRNA) in the mice model and three small interfering RNAs (NR4A3-siRNAs) in the cell culture model. The lipids content, fatty acid synthesis, inflammatory factors, and IR were then assessed with and without NR4A3 knockdown. Furthermore, the underlying mechanism through which NR4A3 exerts its influence was explored by analyzing the interaction between NR4A3 and activating transcription factor 3 (ATF3). Results: In the cell culture experiments, the knockdown of NR4A3 significantly decreased the lipids content, fatty acid synthesis, and inflammatory factors in the LO2 cells treated with FFAs in the NR4A3-shRNA group compared with those in the NC-shRNA control group. In the animal model experiments, NR4A3 knockdown in the HFD male C57BL/6 mice significantly ameliorated HFD-induced hepatic steatosis, inflammation, and IR. Mechanistically, the knockdown of NR4A3 downregulated the expression and transcriptional activity of ATF3, resulting in an impaired ATF3 function. ATF3 overexpression significantly reversed lipid accumulation decline and reduced inflammation after NR4A3 knockdown. Conclusion: The downregulation of NR4A3 alleviates MASLD by modulating ATF3, suggesting this may be a promising therapeutic target.

4
Diversifications of both the three domains of life and SARS-CoV-2 possibly driven by biases between amino acid biosynthetic families

Li, D. J.

2026-06-30 evolutionary biology 10.64898/2026.06.22.733698 medRxiv
Top 0.4%
0.6%
Show abstract

All cellular life forms fall under the three-domain classification of life, raising a fundamental evolutionary question: why does this classification feature three rather than two or four? To answer this question, a more general method, rather than the traditional one based on comparing small-subunit ribosomal RNAs, is required. The three-base periodicity in genomes is a common feature of both cellular life forms and viruses, which is species-specifically biased between amino acid biosynthetic families. Based on comparing such a common feature of all life forms, a global triangular diversification picture has been obtained, whose three angular regions correspond to the three domains, respectively. This mechanism of diversification of life attributes the evolutionary driving forces in diversification of the three domains of life to the biases between amino acid biosynthetic families. Notably, the same mechanism also applies to the contemporary diversification of SARS-CoV-2, whose reasonable results in turn corroborate the above explanation of primordial diversification of life and in addition shed light on the mechanism of speciation.

5
A gapless Landrace pig genome resolves centromeres and telomeres and highlights telomere repeat structures in different pig breeds

Grove, H.; Stenlokk, K. S. R.; Lien, S.; Gjuvsland, A. B.; Arnyasi, M.; van Son, M.; Kent, M.

2026-06-30 genomics 10.64898/2026.06.25.734473 medRxiv
Top 0.4%
0.6%
Show abstract

Abstract The Duroc-derived reference genome Sscrofa11.1 has provided a critical foundation for pig genomics, providing a high-quality reference genome for accurate variant detection and comparative genomics but does not capture breed-specific variation. Here, we present a near-complete, gap-free genome assembly for the Landrace pig (Landrace_v1, GCA_963921485.1), spanning all 20 chromosomes and totaling 2.6 Gb, including 176 Mb of sequence absent from Sscrofa11.1. Comparative analyses with recently published high-quality pig genomes reveal a conserved centromere organization across breeds, accompanied by substantial variation in repeat composition and length, and identify a pig specific pattern of telomere variant repeats across eight pig breeds. The improved resolution of repetitive regions in Landrace_v1 enables more complete reconstruction of complex gene families, including olfactory receptors, and uncovers structural variation at the KIT proto-oncogene receptor tyrosine kinase locus not represented in the Duroc reference. Together, these findings highlight the limitations of single-reference genomes and demonstrate the value of breed-specific assemblies for capturing genomic diversity and improving downstream analyses.

6
Mammalian aging involves genome-wide splicing degeneration leading to functional decline

Zhang, S.;Tyshkovskiy, A.;Ying, K.;Wang, S.;Gladyshev, V.

2026-06-29 Systems Biology 10.64898/2026.06.26.734787 medRxiv
Top 0.4%
0.5%
Show abstract

Alternative splicing exhibits significant changes during development and aging, affecting the composition and variance in the transcriptome. However, it is unclear whether and how age-associated splicing dysregulation leads to functional consequences. Here, an integrative analysis of transcriptome data across mouse and human tissues revealed that aging is characterized by systematic deterioration of the fidelity of RNA splicing, here termed splicing degeneration, a measure of functional alteration of reading frame and domain configuration of protein products. Genes with higher aging-associated splicing degeneration were more conserved and enriched for processes such as RNA metabolism and antigen presentation. By assessing alternative splicing events associated with functional deterioration, we quantified the degree of splicing degeneration. Its level increased with age but was alleviated following calorie restriction or rapamycin treatment, indicating that it can serve as a new molecular hallmark of aging. Mechanistically, through a comprehensive meta-data analysis, we discovered that splicing degeneration is associated with age-associated changes in specific splicing factors, which in turn showed a strong association with age-related transcriptome changes. Overall, our study demonstrates the intricate relationship between aging and genome-wide splicing degeneration, revealing a promising target for aging interventions acting to reverse splicing degeneration.

7
Integrating multi-ancestry common and rare variant mapping accelerates therapeutic target discovery

Koyama, S.; Nakao, T.; Choi, S. H.; Enzan, N.; Jurgens, S. J.; Ellinor, P. T.

2026-07-08 genetic and genomic medicine 10.64898/2026.06.25.26356239 medRxiv
Top 0.6%
0.5%
Show abstract

Integrating human genetics into therapeutic discovery accelerates drug development. However, ancestral biases in historical cohorts have left critical functional variation largely uncharted. Here, we leverage the diverse NIH All of Us Research Program to conduct comprehensive common- and rare-variant association analyses for 624 quantitative traits across 369,655 ancestrally diverse individuals. We identified 6,181 genome-wide significant locus-trait associations (526 novel) and 416 gene-trait associations (105 novel) via rare-variant burden testing. By integrating fine-mapping with computational variant-effect predictors, we systematically prioritized rare, likely causal variants driving these signals. Jointly modeling common and rare variation with protein-class annotations significantly improved the identification of known drug targets compared to common-variant analysis alone. Notably, we identified NRG4 as a high-confidence candidate therapeutic target for preserving kidney function. Our findings demonstrate that characterization of rare and common variation across diverse populations enhances causal gene discovery and identifies novel, actionable therapeutic targets.

8
A Korean pangenome reference of 14 healthy individuals supports structural variant analysis in disease genomes

Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.

2026-07-09 genetic and genomic medicine 10.64898/2026.07.06.26357367 medRxiv
Top 0.6%
0.5%
Show abstract

Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.

9
ThermoFusion: A Multimodal Deep Learning Framework for Generalizable Prediction of Enzyme Thermostability

Wei, Y.; Eberini, I.; Meyer, F.

2026-07-07 bioinformatics 10.64898/2026.07.04.736494 medRxiv
Top 0.8%
0.4%
Show abstract

Protein thermostability is a critical property for both industrial and biomedical enzyme applications, yet experimental evaluation of mutation-induced stability changes remains laborious and costly. Here, we present ThermoFusion, a hybrid deep learning framework that integrates 3D protein structure embeddings from ThermoMPNN with sequence-based embeddings from the pretrained protein language model ESM2 to predict the effects of single-point mutations on protein stability ({Delta}{Delta}G). ThermoFusion exhibits robust generalization, maintaining high predictive accuracy across out of distribution sequences with low identity to the training set -- a scenario where many other machine learning models, including ThermoMPNN and state-of-the-art tools, perform poorly due to reliance on memorization. Benchmarking on a curated enzyme dataset comprising of 105 enzymes and 3144 mutations shows that ThermoFusion reliably identifies stabilizing mutations while accurately predicting stability for enzymes beyond its training set. These results establish ThermoFusion as a powerful tool for rational enzyme design beyond its training set.

10
Brain Structure Shapes Function through higher-order Functional Interactions

Su, S.; Zhuang, M.; Palombo, M.; Liu, M.; Jiang, X.; Zhang, T.; Wang, H.; Zhang, S.

2026-06-28 neuroscience 10.64898/2026.06.22.733911 medRxiv
Top 0.8%
0.4%
Show abstract

Brain function is deeply embedded within multiscale structural architecture. Conventional studies predominantly utilize pairwise connectivity networks to investigate structure-function relationships. However, this low-dimensional perspective overlooks multi-region collaborations for complex cognition. Consequently, whether and how anatomy constrains such higher-order functional networks remains unresolved. To address this pivotal question, we utilize an information-theoretic O-information approach to characterize higher-order functional interactions (HOIs). By reconstructing individual-level HOIs from multimodal structural networks, we directly validate the structural constraint on HOIs. The resulting reconstruction coefficients are defined as structural-functional constraint strength (SFCS), serving as a quantitative vehicle to decipher how anatomy shapes these higher-order networks. SFCS uncovers a highly heterogeneous structural constraint landscape across data modalities, spatial regions, and informational interaction modes. Crucially, individualized SFCS robustly predicts multi-domain cognitive phenotypes, showing higher sensitivity for informant-reported than patient-reported assessments. Finally, we show that this landscape undergoes pathological, mode-specific reorganization in Alzheimers disease. Cross-scale alignment with spatial transcriptomics further demonstrates that this macroscale network remodeling is coupled with microscale metabolic and regulatory gene pathways. Collectively, our findings not only validate the structural constraint on higher-order functional networks but also decipher its precise underlying mechanisms. This constraint paradigm plays a pivotal role in shaping diverse cognitive capabilities, while its pathological disruption in Alzheimers disease highlights the potential of SFCS as a biomarker for tracking neurodegenerative network impairments.

11
Retina-derived Quantitative Biomarkers of Brain Health

Ma, T.; Yan, T.; Sun, J.; Wu, N.; Xu, M.; Zhang, R.; Zeng, N.; Sun, Q.; Hui, Y.; Wu, Y.; Wang, Z.; Wong, T. Y.; Lv, H.; Qiao, H.

2026-07-08 public and global health 10.64898/2026.07.05.26357344 medRxiv
Top 0.8%
0.4%
Show abstract

Accurate and scalable assessment of quantitative neuroimaging biomarkers, such as white matter hyperintensities (WMH) and hippocampal (HIP) volumes, is essential for understanding and monitoring brain health, preventing neurological diseases and improving healthspan. However, population-level evaluation of these neuroimaging biomarkers relies on inaccessible, costly and time-consuming magnetic resonance imaging (MRI). Here we propose RetiBrain, a cross-modal deep learning framework that predicts these neuroimaging biomarkers from retinal color fundus photography (CFP) images. By distilling latent structural representations from MRI-based models into a CFP-based model, RetiBrain establishes biologically grounded eye-to-brain mapping. In a CFP-MRI paired cohort, RetiBrain accurately estimates six WMH- and HIP-related biomarkers and outperforms the state-of-the-art retinal foundation model RETFound, improving the mean Pearson correlation coefficient by 0.309 (from 0.240 to 0.549) and achieving a coefficient of 0.640 for periventricular WMH prediction. By integrating structural, topological and geometric feature analyses from CFP images, RetiBrain identifies interpretable retinal representations associated with neurodegeneration and cerebrovascular injury, hallmarks of major neurological diseases such as dementia and stroke. In a longitudinal cohort comprising 2,082 participants (4,164 CFP images with up to 15 years of follow-up), RetiBrain-predicted neuroimaging biomarkers robustly estimated neurological disease risk, as illustrated by dementia prediction (AUROC of 0.824, hazard ratio 2.500 per standard deviation increase, 95% CI: 2.201-2.840). RetiBrain provides a robust, scalable, cost-effective and convenient approach for the assessment of neuroimaging biomarkers, and has potential for long-term brain health monitoring in large-scale general population settings.

12
Deoxyribonucleotide incorporation reshapes mRNA design beyond canonical ribonucleotide boundaries

Ding, X.; Liao, R.; Bampi, G. B.; Zhang, D.; Guan, S.; Rosenecker, J.

2026-07-10 synthetic biology 10.64898/2026.07.09.737403 medRxiv
Top 0.8%
0.4%
Show abstract

Messenger RNA (mRNA) is canonically composed of ribonucleotides, with sporadic incorporation of deoxyribonucleotides into natural RNA transcripts being traditionally regarded as a rare, deleterious error arising from transcriptional infidelity. Here, we challenge this paradigm by demonstrating controlled partial substitution of ribonucleotides with deoxyribonucleotides during in vitro transcription (IVT) generates intact, stable and fully translationally competent IVT-mRNA. Unexpectedly, chimeric DNA-RNA backbone modification exhibits markedly enhanced IVT-mRNA translation several fold across multiple cell types and in vivo via diverse dosing routes relative to their ribonucleotide-based counterparts. 25% substitution of cytidine triphosphate with deoxycytidine triphosphate achieved best-performing translational output, surpassing the current gold-standard N1-methylpseudouridine (m1{Psi})-modified IVT-mRNA in a B16-OVA tumor vaccination model. These findings identify nucleotide class composition as a previously unrecognized parameter governing IVT-mRNA function and establish hybrid ribonucleotide-deoxyribonucleotide backbone engineering as a versatile strategy to expand the chemical space for next-generation mRNA therapeutics.

13
Directed evolution of compact synthetic promoters via AlphaGenome and genetic algorithms

Nie, L.

2026-07-09 synthetic biology 10.64898/2026.06.28.735069 medRxiv
Top 0.8%
0.4%
Show abstract

Compact tissue-specific promoters are highly desirable for gene therapy because viral vectors possess limited packaging capacity. However, existing promoter engineering strategies rely primarily on rational design or de novo sequence generation and lack efficient approaches for compressing long native promoters while preserving regulatory specificity. Although genome foundation models have substantially improved sequence-to-function prediction, they have not been effectively translated into computational platforms for promoter engineering. Here, we present VirEvo, a computational promoter engineering framework that integrates a virtual dual-luciferase assay (VirDLA), genome-foundation-model-guided genetic evolution, and an orthogonal Pan-Tissue Consistency Filter (PTCF). VirDLA introduces an internal-reference normalization strategy inspired by dual-luciferase reporter assays, enabling relative comparison of promoter activity across tissues without retraining AlphaGenome. Guided by these normalized activity scores, VirEvo iteratively optimizes promoter selectivity, off-target activity, and sequence length. Using the human p16INK4a promoter as a proof of concept, VirEvo evolved a compact synthetic promoter, SRP2M, of only 398 bp, representing an 85.9% reduction in sequence length. Experimental validation using dual-luciferase reporter assays in senescent IMR90 fibroblasts demonstrated that SRP2M retained 77% of wild-type senescence selectivity while reducing basal leakage to 52% of the wild-type level. Together, these results demonstrate the feasibility of genome-foundation-model-guided promoter engineering. VirEvo provides a generalizable framework for designing compact tissue-specific regulatory elements and extends the application of genome foundation models from functional prediction to synthetic regulatory engineering.

14
Charge-trap flash memory cells of the brain

Foster, P. P.; Chhikara, R. S.; Boriek, A. M.

2026-07-03 neuroscience 10.64898/2026.06.29.733154 medRxiv
Top 0.9%
0.3%
Show abstract

Despite extensive study of cellular mechanisms underlying long-term potentiation, no single specific protein or gene has been identified which encodes an individual unit of information, or memory bit. Indeed, the brain engram remains a knowledge gap. The theory of exclusion led us to cancel one-by-one several unrealistic biological options, suggesting that the explanation resides somewhere else. Superposition of up to concentric 300 myelin layers, spiraled, and highly compacted wrapping a single axon and each wrap could host hundreds to thousands of niches, as memory cells, collectively consisting of a massive array of cells. The disjointed 3D spatial superposition allows storage of charges, nodes not facing from a layer to next. The thickness of a single myelin layer ranges from 7.0 to 20 nm. The dimension scale is approximately the exact dimensions of the charge trap, the tunnel and dielectric also equipping current AI microchips. Stored charges are positive ions, with similar effect whether charges are negative or positive charges creating an electromagnetic field. To write data, following an action potential, this voltage applies to the control gates of the myelin layers producing an ionic charge injection. This causes charges to gain energy and tunnel through the myelin layer across Ranvier nodes, via quantum tunneling, and deep into the concentric myelin multilayers. This is creating an insulated trapping of K+ ions isolated from the system. In a long white matter tract bundle, the near-perfect isolation of millions of axons within compressed myelin wrap-ion channel K+/Na+ systems provides quantum coherence and precision of asynchronous firing property. The injected ionic charges (K+) become physically stuck in traps within the myelin layers. The K+ ions may not move freely, completely trapped after AP ceases. Mirroring a single-bit, single-level-cell, a trapped ionic charge (ions K+) may represent a 1, while an empty cell (absence of K+) represents a 0. The trial-and-error process, with a Bayesian inference which may have also been the core evolution of the learning human brain. Based on selected mathematical equations, we analyzed the general scheme on how deep learning may be embedded in the brain

15
CellDF: Quality-controlled cell matching for whole-slide HE-IHC label transfer

Jang, E.; Huh, Y.-M.

2026-06-24 pathology 10.64898/2026.06.18.733058 medRxiv
Top 0.9%
0.3%
Show abstract

Serial-section immunohistochemistry (IHC) is the largest available source of paired hematoxylin and eosin (HE) and IHC whole slide images, yet it remains underexploited for cell-level supervision: adjacent sections sample non-identical cells, and residual registration error prevents direct assignment of IHC labels to individual HE cells. We present CellDF (Cell Displacement Field), which turns registered serial-section data into pairs of HE cells and their IHC labels by solving cell matching at whole-slide scale and assessing its reliability without ground-truth correspondences. CellDF estimates a locally adaptive residual displacement field through iterated kernel regression over each HE cells K nearest IHC candidates; a sparse-kernel variant keeps it tractable at the cell counts of a whole slide, where pairwise matchers are not. The within-tile distribution of the estimated displacements yields two ground-truth-free statistics, the directional scatter{sigma}{theta} and the between-tile angular deviation |{Delta}{theta}|, that localize matching quality more finely than landmark-based target registration error and drive a two-stage outlier filter that withholds labels where matching is unreliable. On 54 same-section HyReCo pairs,{sigma}{theta} correlates only moderately with landmark error and flags localized restaining damage that global error misses; on 30 four-marker Acrobat serial-section cases, the same statistic flags which IHC marker, if any, lies physically close enough to HE to support cell-level transfer. As a proof of concept, IHC labels transferred through CellDF trained a cell classifier on HE embeddings that generalized to held-out cells within the sample (F1 0.85, AUROC 0.88), establishing serial-section IHC as a usable cell-level labeling resource. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=78 SRC="FIGDIR/small/733058v1_ufig1.gif" ALT="Figure 1"> View larger version (42K): org.highwire.dtl.DTLVardef@a9b3dcorg.highwire.dtl.DTLVardef@15f652corg.highwire.dtl.DTLVardef@1eb3396org.highwire.dtl.DTLVardef@87dda2_HPS_FORMAT_FIGEXP M_FIG C_FIG

16
Generative Drug Design in a Loop with dtSFM

Reddy, S. T.

2026-07-08 synthetic biology 10.64898/2026.06.10.731501 medRxiv
Top 0.9%
0.3%
Show abstract

Directed evolution consisting of iterative rounds of diversification, selection, and counter-selection, underlies modern protein and antibody engineering, yet small-molecule drug design still advances largely through high-throughput screening and medicinal-chemistry intuition. Transformer softmax attention is mathematically identical to the Boltzmann distribution that governs molecular binding at thermal equilibrium1, an isomorphism that prescribes a sequence-native Specificity Foundation Model (SFM)2. This framework was recently applied across seven molecular recognition domains3,4 and scaled into the drug-target SFM (dtSFM), the first to pair a full-scale encoder with a generative decoder5. Whether such a model can be driven, iteratively and under selection, to optimize leads rather than sample them once has not been shown. Here we present GenLoop, a closed generative drug design loop that turns single-pass generation into directed evolution of chemistry. dtSFM generates target-conditioned molecules and reranks them by their thermodynamic compatibility score. An orthogonal structural verifier, AlphaFold 3, is used that shares no architecture or training data with dtSFM. Cheminformatics filters enforce developability, and generative evolution is performed on the structurally verified candidates, selecting for predicted binders and counter-selecting against off-target chemistry. Applied across twelve drug targets spanning pharmacologically distinct mechanism classes, GenLoop produced AlphaFold 3-verified designs that reached the structural confidence of the approved drug for five of the twelve targets, with the best designs at interface iPTM 0.93-0.98 and PAE 0.8-2.0 [A], as well as resolving paralog selectivity across nine targets. Two full disease campaigns followed. For the cystic-fibrosis transmembrane conductance regulator, GenLoop designed nine developability-filtered and structurally novel lead candidates (iPTM up to 0.93, interface PAE 2.3 [A]) targeting all three orthogonal sites of the approved drug Trikafta. For the GLP-1 receptor family, dtSFM engineered tunable single-, dual-, and triple-receptor incretin designs, yielding 23 central-pocket candidates that are structurally novel at median iPTM 0.89 and interface PAE 1.95 [A]. GenLoop with dtSFM brings directed evolution to small molecules through computational-thermodynamic selection; wet-lab validation is the immediate next step.

17
Rett syndrome lifespan extension in mice via AI-guided ADAR editing

Savva, Y. A.; Booth, B. J.; Shumaker, L.; Fasnacht, R.; Burleigh, S. M.; Jiang, Y.; Cao, Y.; Johnson, B.; Bagepalli, L. R.; Golic, F.; Enger, N.; Feiring, R.; Sadowski, A.; Rich, S.; Lakshmanan, A.; Milani, N.; Chadwick, E. M.; Hauskins, C.; Works, M. G.; Huss, D. J.; Briggs, A. W.; VanSchoiack, A. A.

2026-06-28 neuroscience 10.64898/2026.06.23.734060 medRxiv
Top 1%
0.3%
Show abstract

Rett syndrome is a severe neurodevelopmental disorder primarily caused by mutations in the MECP2 gene. A significant subset of severe cases are driven by nonsense mutations that generate premature stop codons, leading to loss of functional MeCP2 protein. Here, we describe a novel therapeutic strategy that uses endogenous adenosine deaminase acting on RNA (ADAR) enzymes to correct the R168X mutation at the RNA level. Using generative artificial intelligence trained on large empirical datasets, we engineered guide RNAs (gRNAs) that recruit endogenous ADAR to convert the mutant stop codon (UGA) into a tryptophan (UGG) to restore full-length MeCP2. Once incorporated into an optimized expression system based on endogenous small nuclear RNA regulatory elements and packaged into adeno-associated virus, these gRNAs enabled precise RNA editing at the target site with minimal off-target activity across the transcriptome while restoring full-length MeCP2 protein expression in patient-derived induced pluripotent stem cell neurons. Delivered intravenously to an R168X mouse model, the gRNAs achieved ~70% targeted RNA editing and substantially restored MeCP2 throughout the brain resulting in markedly improved Rett-like phenotypes and significantly extended lifespan. These findings demonstrate that AI-guided ADAR-mediated RNA editing is a precise and efficient technology for correcting nonsense mutations, with therapeutic potential for Rett syndrome and other genetic diseases.

18
AI-enabled rhodopsin design for blue-light enhanced bacterial growth

Saeed, H.;Lewis, M.;Fujiwara, T.;Huang, J.;Konno, M.;Mori, K.;Yoshizawa, S.;Inoue, K.;Pan, T.;Wang, Y.;Yang, A.;Huang, W.

2026-06-30 Synthetic Biology 10.64898/2026.06.29.735265 medRxiv
Top 1%
0.3%
Show abstract

We developed an AI-guided design pipeline that generated and validated non-natural microbial rhodopsins with spectral properties not yet known in nature. The pipeline comprised a three-stage in silico design, a genetic algorithm (GA) for sequence generation, a stacked LASSO and XGBoost machine-learning (ML) regressor for spectral prediction and fitness ranking, and a Markov-based sequence plausibility filter to enforce proton pumping like characteristics. Four candidate rhodopsins (APR1, APR2, APR6, and APR7) targeting blue light absorption were designed and AlphaFold3 structural modelling predicted retinal binding pocket architecture consistent with outward proton-pumping function. Experimental characterisation confirmed that all four variants absorbed light at [~]410 nm and significantly promoted the growth of Cupriavidus necator under blue light illumination. This study demonstrates that AI-enabled design can engineer proteins with no natural precedent, generating light-harvesting rhodopsins with novel spectral properties while preserving biological function, marking a significant advance in programmable synthetic biology.

19
Lineage-Specific Neofunctionalization of Polyacetylene-Directing UGT76 Glycosyltransferases in Campanulaceae

Qiu, S.; Hu, J.; Cao, X.; He, M.; Wang, C.; Di, P.; Chen, S.; Zhang, C.; Xiao, Y.; Mao, R.; Sun, W.; Chen, W.

2026-07-03 evolutionary biology 10.64898/2026.07.02.735953 medRxiv
Top 1%
0.3%
Show abstract

Polyacetylene glycosides exhibit notable pharmacological activities, yet the glycosyltransferases acting on their polyacetylene scaffolds remain unknown. Here we report a telomere-to-telomere genome assembly of Codonopsis pilosula and, guided by spatial metabolomics, characterize three UDP-glycosyltransferases: CpUGT76BG1 and CpUGT76BG2 catalyze the direct glycosylation of lobetyol to lobetyolin, while CpUGT94BY2 performs subsequent sugar-sugar coupling to produce lobetyolinin, with each activity confirmed by in planta overexpression. Structural modeling reveals that CpUGT76BG1 and CpUGT76BG2 employ a deep hydrophobic tunnel to fully encase the linear polyacetylene chain, a binding architecture distinct from the shallow pockets used by canonical plant UGTs for planar aromatic substrates. Ancestral sequence reconstruction across eleven nodes partitions the UGT76 lineage into three functionally distinct evolutionary stages, tracing the trajectory from an ancestral shallow pocket to this specialized deep architecture. These findings establish the key glycosylation steps of polyacetylene glycoside biosynthesis, define a tunnel-based paradigm for non-planar substrate recognition, and reveal how tandem duplication-driven active site remodeling generates metabolic novelty.

20
ComplexDesign: sequence-hallucination design of protein binders bridging multiple proteins

Xu, J.; Ren, M.; Qi, N.; Zhang, X.; He, Z.; Yu, C.; Bu, D.

2026-06-24 bioinformatics 10.64898/2026.06.21.733655 medRxiv
Top 1%
0.3%
Show abstract

MotivationDesigning multichain protein complexes requires coordinating the folding of component proteins with the formation of their interfaces. The existing methods, however, remain limited in their ability to satisfy these requirements simultaneously, especially for trimeric and tetrameric complexes. As an important practical scenario, designing a binder that bridges two target proteins into a ternary complex requires flexibility in the relative arrangement of the two targets, adding an additional challenge to existing design methods. ResultsWe present ComplexDesign, a hallucination-based approach for multichain protein design. ComplexDesign performs structure-prediction-guided sequence optimization to simultaneously fold each protein chain and form inter-chain interactions that bind them together. To provide the flexibility required to appropriately arrange these target proteins, ComplexDesign introduces a specialized masking mechanism that enables exploration of possible relative arrangements rather than being limited to the predefined ones. Across a comprehensive set of benchmarks with various chain lengths, ComplexDesign outperformed existing methods in the unconditional design of dimers, trimers, and tetramers, achieving a high design success rate exceeding 50%, supporting its capability for multichain complex design. Furthermore, in the case of multi-target binder design, ComplexDesign produced high-confidence, self-consistent ternary complexes for 8 out of 10 target pairs. These results establish ComplexDesign as an effective tool for multichain protein design, with particular utility for designing binders that bridge two target proteins. Availability and implementationThe source code of ComplexDesign will be made publicly available upon publication.